Papers with English corpus

23 papers
Understanding the Use of Quantifiers in Mandarin (2022.findings-aacl)

Copied to clipboard

Challenge: a corpus of short texts in Mandarin is analyzed to examine the "coolness" hypothesis . quantified expressions are used to describe short texts, but are not as informative as English .
Approach: They propose a corpus of Mandarin in which quantified expressions figure prominently.
Outcome: The proposed corpus of short texts in Mandarin is compared with an English corpus.
How to Translate Your Samples and Choose Your Shots? Analyzing Translate-train & Few-shot Cross-lingual Transfer (2022.findings-naacl)

Copied to clipboard

Challenge: Recent studies have focused on zero-shot cross-lingual transfer of pretrained languages.
Approach: They propose to use few-shot cross-lingual transfer to improve zero-shot performance of multilingual pretrained language models.
Outcome: The proposed model can be scaled to high-quality samples and improves on zero-shot performance.
Probe-Less Probing of BERT’s Layer-Wise Linguistic Knowledge with Masked Word Prediction (2022.naacl-srw)

Copied to clipboard

Challenge: Among studies on localization of linguistic knowledge, it is unclear what information is encoded in each layer.
Approach: They analyze BERT’s layer-wise masked word prediction on an English corpus and find syntactic and semantic information is encoded at different layers for words of different syntaktic categories.
Outcome: The proposed model outperforms state-of-the-art models in many downstream tasks.
A Model of Cross-Lingual Knowledge-Grounded Response Generation for Open-Domain Dialogue Systems (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on open-domain dialogue systems that allow free topics are challenging . however, non-English dialogue systems suffer from reproducing the performance of English dialogue systems .
Approach: They propose to use English knowledge to improve the performance of open-domain dialogue systems . they construct a Korean-English T5 language model and develop a knowledge-grounded Korean dialogue model .
Outcome: The proposed model improves even when only English knowledge is given . the model is built with a pre-trained language model and a knowledge-grounded Korean dialogue model .
MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing zero-shot dialogue generation systems rely on large-scale pre-trained language models.
Approach: They propose a multilingual learning framework for zero-shot dialogue generation that can transfer knowledge from an English corpus to a non-English corpus with zero samples.
Outcome: The proposed framework can transfer knowledge from an English corpus to a non-English corpus with zero samples.
OntoGUM: Evaluating Contextualized SOTA Coreference Resolution on 12 More Genres (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for coreference resolution are unable to evaluate generalizability to open domain data.
Approach: They propose to make an OntoNotes-like coreference dataset publicly available and convert it into an English corpus.
Outcome: The proposed dataset is the largest human-annotated coreference corpus following the OntoNotes guidelines and the first to be evaluated for consistency with the OnToNote's scheme.
Predicting emergent linguistic compositions through time: Syntactic frame extension via multimodal chaining (2021.emnlp-main)

Copied to clipboard

Challenge: Natural language relies on a finite lexicon to express an unbounded set of emerging ideas.
Approach: They propose a framework that exploits the cognitive mechanisms of chaining and multimodal knowledge to predict emergent compositional expressions through time.
Outcome: The proposed framework predicts emergent compositions through time using cognitive mechanisms . it is based on modal knowledge and categorizing models of chaining in a syntactically parsed English corpus .
El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing (2021.eacl-main)

Copied to clipboard

Challenge: Code-switching (CS) is the alternation of languages within an utterance or conversation.
Approach: They propose to use translation-and-align and augment with a generation model followed by match-and filter to improve CS generalizability of cross-lingual models when data for only one language is available.
Outcome: The proposed models improve when only English data is available alongside zero or a few CS training instances.
KTH Tangrams: A Dataset for Research on Alignment and Conceptual Pacts in Task-Oriented Dialogue (L18-1)

Copied to clipboard

Challenge: Existing studies on instructor-manipulator dialogue use disparate but similar datasets . a recent study examined the alignment of referring expressions (RL) in situated dialogue .
Approach: They propose to use a corpus of referring expressions in a relatively free dialogue with physical features generated in simulated situations to study alignment in referring language.
Outcome: The proposed datasets facilitate analysis of dialogic linguistic phenomena regarding alignment in the formation of referring expressions known as conceptual pacts.
Exploiting Entity BIO Tag Embeddings and Multi-task Learning for Relation Extraction with Imbalanced Data (P19-1)

Copied to clipboard

Challenge: Existing methods to perform relation extraction are feature-based or kernel-based, but the results of our study show that they can improve the performance of a baseline model with more than 10% absolute increase in F1-score.
Approach: They propose a multi-task architecture which jointly trains a model to perform relation identification with cross-entropy loss and relation classification with ranking loss.
Outcome: The proposed model outperforms the state-of-the-art models on ACE 2005 Chinese and English corpus and significantly improves the performance of a baseline model with more than 10% increase in F1-score.
x-enVENT: A Corpus of Event Descriptions with Experiencer-specific Emotion and Appraisal Annotations (2022.lrec-1)

Copied to clipboard

Challenge: Emotion classification is often formulated as the task to categorize texts into a predefined set of emotion classes.
Approach: They propose that a classification setup for emotion analysis should be performed in an integrated manner, including the different semantic roles that participate in an emotion episode.
Outcome: The proposed method reveals patterns in the co-occurrence of people’s emotions in interaction.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
Semantic Supersenses for English Possessives (L18-1)

Copied to clipboard

Challenge: Existing semantic categories for possessive constructions are limited to nominals and s-genitives.
Approach: They propose to use a supersense inventory to annotate English possessives . they show existing supersensor categories are readily applicable to possessives.
Outcome: The proposed annotations are applied to English possessives in a corpus of web reviews.
On Difficulties of Cross-Lingual Transfer with Order Differences: A Case Study on Dependency Parsing (N19-1)

Copied to clipboard

Challenge: Existing studies on crosslingual transfer have focused on word-level information sharing, but words are not independent in sentences; their combinations form larger linguistic units, known as context.
Approach: They propose to use orderagnostic models to transfer word order to distant languages . they train dependency parsers on an English corpus and evaluate their transfer performance on 30 other languages.
Outcome: The proposed model performs better on languages with different word orders than on other languages.
GenWebNovel: A Genre-oriented Corpus of Entities in Chinese Web Novels (2025.coling-main)

Copied to clipboard

Challenge: Existing literature on nested entity recognition is insufficient partly due to insufficient annotated data.
Approach: They propose a method that utilizes a pre-trained language model as an In-context learning example retriever to boost the performance of large language models.
Outcome: The proposed method significantly enhances entity recognition, matching state-of-the-art (SOTA) models without additional training data.
The Nunavut Hansard Inuktitut–English Parallel Corpus 3.0 with Preliminary Machine Translation Results (2020.lrec-1)

Copied to clipboard

Challenge: Inuktitut language is a member of the Inuit-Yupik-Unangan family . it is spoken in two territories, Nunavut and the Northwest Territories .
Approach: They describe a sentence-aligned Inuktitut–English corpus released in Nunavut . it is the largest parallel corpus of a polysynthetic language released to date . they also describe preliminary experiments on machine translation between the languages .
Outcome: The proposed corpus is the largest sentence-aligned corpus of a polysynthetic language or an Indigenous language of the Americas . the alignments were evaluated and the results were compared with other methods .
Intent Classification and Slot Filling for Privacy Policies (2021.acl-long)

Copied to clipboard

Challenge: Sentences written in privacy policies explain privacy practices and the constituent text spans convey further specific information.
Approach: They propose an English corpus of 5,250 intent and 11,788 slot annotations . they propose two alternative neural approaches to model the corpus as a sequence-to-sequence learning task.
Outcome: The proposed corpus predicts intent classification and slot filling, while the sequence tagging method outperforms slot filler by a large margin.
Discovering the Language of Wine Reviews: A Text Mining Account (L18-1)

Copied to clipboard

Challenge: odors and flavors are often expressed in wine reviews, but they are often not.
Approach: They use a corpus of wine reviews to find out what wine is like in a review . they use lexical bag-of-words features, domain-specific terminology features and word embedding features to train machine learning.
Outcome: The proposed model predicts the wine's color, grape variety, and country of origin based on the review text alone.
Transformer-based Speech Model Learns Well as Infants and Encodes Abstractions through Exemplars in the Poverty of the Stimulus Environment (2025.coling-main)

Copied to clipboard

Challenge: Existing theories of language learning for infants are inadequate, according to Chomsky . infants learn language in impoverished environments, according a new study .
Approach: They designed a series of tasks, scenarios, and metrics to simulate the POS . they found that the emerging speech model wav2vec2.0 can learn well in noisy Mandarin environments.
Outcome: The proposed model can learn in noisy and sparse Mandarin environments.
Mining Crowdsourcing Problems from Discussion Forums of Workers (2020.coling-main)

Copied to clipboard

Challenge: Among the most widely used platforms are Upwork, Appen, and above all Amazon Mechanical Turk (MTurk) which host annotation tasks and collect huge sets of annotated data from workers.
Approach: They propose to use topic modeling to analyze workers' complaints from a new English corpus of workers’ forum discussions to identify problems in task design, task operation, and task evaluation that workers face with requesters in crowdsourcing processes.
Outcome: The findings form the basis for future research on how to improve crowdsourcing processes.
ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Abuse Detection in Conversational AI (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on abusive language towards conversational AI systems are not conclusive as they are not performed with live systems nor with real users due to the lack of reliable abuse detection tools.
Approach: They propose to use a convAI dataset to account for the complexity of the task and to bench-mark existing models against this data.
Outcome: The proposed model shows that abuse distribution is different compared to other datasets, with sexual tinted aggression towards the virtual persona of the systems.
Language Proficiency Scoring (2020.lrec-1)

Copied to clipboard

Challenge: a new paper evaluates and extends the results of an automated proficiency classification system for different languages.
Approach: They propose to extend an automated essay scoring system proposed by CEFR . they compare results with those from previous paper and add a new corpus for english .
Outcome: The proposed approach does not scale well with the added English corpus.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations